Skip to content

ci: bump base image/runner to v26.4 and drop v26.3-only workarounds#896

Open
WangLingxun wants to merge 9 commits into
mainfrom
chore/ci-bump-primus-v26.4
Open

ci: bump base image/runner to v26.4 and drop v26.3-only workarounds#896
WangLingxun wants to merge 9 commits into
mainfrom
chore/ci-bump-primus-v26.4

Conversation

@WangLingxun

@WangLingxun WangLingxun commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

Upgrade the Primus runtime image from rocm/primus:v26.3 to v26.4 across CI Dockerfile/workflow, example launch scripts, benchmark scripts and docs, drop v26.3-only CI workarounds now that they're no longer needed, and fix several bugs found while validating the change end-to-end. JAX/MaxText base image not upgraded this round.

Image bump

  • .github/workflows/ci.yaml & docker/Dockerfile: BASE_IMAGE -> v26.4, sync v26.3 comments
  • README/examples/docs/benchmark scripts: bump documented image tag to v26.4, release branch to release/v26.4, pip pin to primus==26.4.0
  • runs-on stays on the v26.3 self-hosted runner label (no v26.4 runner registered yet)

Workaround cleanup

  • Drop the v26.3-only "Purge stale TorchTitan install" step: v26.4 no longer bundles a torchtitan that shadows the submodule.
  • Drop the origami (rocm-libraries@223648a) install from Dockerfile and CI: this base image ships no origami, so primus_turbo falls back to heuristic kernel selection for MoE grouped-gemm. Left the pip install one-liner in a comment to add it back if ever needed.
  • aiter_deepbind_patches.py self-detects whether the RTLD_DEEPBIND workaround is still needed: skips installing it once transformer_engine.__version__ >= 2.14 (fixed in v26.4), still applies on v26.3 and older. Added unit tests for the gating logic.
  • Drop the from-source rocSHMEM rebuild and the v26.3 ROCM_HOME/UCX_HOME/MPI_HOME overrides: v26.4 already ships a working prebuilt rocSHMEM and correct env vars for its pip-installed ROCm SDK layout; the old overrides pointed at a now-nonexistent /opt/rocm and broke the rebuild.

v26.4 pip-installed ROCm SDK fallout

v26.4 switched from a filesystem-installed /opt/rocm to a pip-installed rocm-sdk-devel venv package. Several things assumed the old layout:

  • Dockerfile.ainic: ROCM_PATH/MPI_PATH were hardcoded to v26.3 paths; rccl/amd-anp also hardcode /opt/rocm internally, and amdclang++ couldn't find its sibling clang++ (moved under llvm/bin/). Fixed via symlinks instead of patching those repos, plus added pipefail so cmd | tee log steps fail on cmd's exit code (was silently swallowing failures until a much later, unrelated-looking step).
  • torch is now a proper pip wheel with a strict Requires-Dist: triton==<exact commit> in its metadata. Our Dockerfile force-reinstalls a different pinned triton commit without updating that metadata, so any later pip install (e.g. the unit-test job's requirements.txt) saw torch's declared triton dependency as unsatisfiable and silently replaced torch with a plain PyPI/CUDA build, breaking torchvision and other ROCm extensions. Fixed by patching torch's metadata to the actually-installed triton version right after building it. Also restored the TRITON_COMMIT build-arg, which had been dead (hardcoded to 09500db9 in the Dockerfile) since Build Triton from source and bump Primus-Turbo  #694.

CI fix

  • The unit-test container reuse check compared only the image tag string (docker.io/tasimage/primus:${IMAGE_TAG}), not content. On main, IMAGE_TAG is always latest, so after build-docker repushes under that tag, the check still saw an identical string and reused the stale container instead of pulling the new image. Fixed by pulling first and comparing the resolved image id instead of the tag name.

@WangLingxun
WangLingxun marked this pull request as draft July 21, 2026 06:21
- ci.yaml: BASE_IMAGE + runs-on v26.3 -> v26.4; drop the v26.3-only
  'Purge stale TorchTitan' step (v26.4 ships no torchtitan to shadow).
  Keep 'Install fixed origami' (v26.4 still ships no origami).
- Dockerfile: default BASE_IMAGE + comments v26.3 -> v26.4.
- docs/scripts: bump default image refs to v26.4, release branch to
  release/v26.4, and pip pin to primus==26.4.0.
@WangLingxun
WangLingxun force-pushed the chore/ci-bump-primus-v26.4 branch 2 times, most recently from a6e9036 to 90e0389 Compare July 21, 2026 09:24
- Dockerfile/ci.yaml: drop the origami install; base image ships none,
  so primus_turbo falls back to heuristic MoE grouped-gemm kernel
  selection.

- aiter_deepbind_patches.py: skip the RTLD_DEEPBIND hook once
  transformer_engine >= 2.14 (fixed in rocm/primus:v26.4); still
  applies on older images.

- Add unit tests for the new gating logic.
Container reuse compared only the image tag string, not content. On
main, IMAGE_TAG is always "latest", so after build-docker repushes
under that tag, the check still saw an identical string and reused the
stale container instead of pulling the new image.

Fix: pull first, then compare docker image inspect .Id (target) against
the container's docker inspect .Image (actual) instead of the tag name.
@WangLingxun
WangLingxun force-pushed the chore/ci-bump-primus-v26.4 branch from 62f555d to 11e1885 Compare July 23, 2026 10:49
@WangLingxun
WangLingxun marked this pull request as ready for review July 23, 2026 10:54
v26.4 ships a prebuilt rocSHMEM plus correct ROCM_HOME/UCX_HOME/MPI_HOME
for its pip-installed ROCm SDK layout. Rebuilding rocSHMEM from source
and re-pinning those vars to their old v26.3 paths made the rebuild
CMake step fail (ROCM_HOME pointed at a now-nonexistent /opt/rocm) and
is redundant anyway, so drop both. The restore instructions left in the
Dockerfile intentionally read the vars from the base image instead of
re-declaring them, so they stay correct across future base image bumps.
@WangLingxun
WangLingxun force-pushed the chore/ci-bump-primus-v26.4 branch from 7629d86 to 1a48a5b Compare July 24, 2026 07:56
Same root cause as the earlier Dockerfile fix, in a file that hadn't been
touched yet: ROCM_PATH/MPI_PATH were hardcoded to v26.3's /opt/rocm and
system-OpenMPI paths, which no longer exist under v26.4's pip-installed
rocm-sdk-devel. rccl/amd-anp also hardcode /opt/rocm internally, and
amdclang++ can't find its sibling clang++ (moved under llvm/bin/), so
symlink both in rather than patching those repos. Also add pipefail so
`cmd | tee log` steps fail on `cmd`'s exit code instead of tee's, which
was silently swallowing the above failures until a much later, unrelated-
looking step. Verified locally end to end (rccl + amd-anp build install).
Our custom triton build (pinned commit) leaves torch's dist-info
Requires-Dist pointing at the base image's stock triton, so later
pip installs (e.g. unit-test CI) see it as unsatisfiable and silently
replace torch with a plain PyPI/CUDA build. Patch the metadata to the
actual triton version right after installing it, and restore the
TRITON_COMMIT build-arg that had been dead since #694.
…symbol

Primus-Turbo links librocshmem.a via -l: (not --whole-archive) under
-fgpu-rdc/--hip-link, so rocSHMEM's team.cpp.o - which defines the plain
rocshmem::ROCSHMEM_TEAM_WORLD symbol - never gets pulled into
libprimus_turbo_kernels.so. Being a shared object, the missing symbol
doesn't fail the build; it only surfaces as an ImportError at runtime.
Fix by re-exporting that symbol from a small shim .so and wiring it in
via patchelf.
Under PyTorch 2.12's AOTAutograd, letting Dynamo trace into this
Function's forward (the default for the modern forward/setup_context
split) silently computes wrong gradients for the training-only proxy
enabled by tests/unit_tests/.../test_fp8_compile_graph_breaks.py
(TestDualFP8Convergence::test_dual_compiled_setup_ctx_vs_eager passed
under v26.3/torch 2.10, diverged to rel_diff~0.2-0.3 under v26.4/torch
2.12). Reproduced on real MI300X hardware with backend="eager" (pure
Dynamo graph capture, no Inductor codegen) showing the identical
divergence, which rules out kernel fusion/codegen and points at
Dynamo/AOTAutograd's handling of the 6 forward outputs that exist only
to be captured by setup_context/save_for_backward and are otherwise
unused by the outer graph. Matches pytorch/pytorch#131794, apparently
regressed by the AOTAutograd unused-output backward-pruning rework in
pytorch/pytorch#186355.

@torch._dynamo.allow_in_graph makes Dynamo treat .apply() as a single
opaque node instead of inlining into forward, sidestepping the bug
while still keeping everything in one fused graph (verified via
torch._dynamo.explain: graph_count=1, graph_break_count=0). Same
mechanism the file's own @allow_in_graph baseline classes already use.

Verified locally on MI300X: the full
tests/unit_tests/backends/megatron/diffusion/test_fp8_compile_graph_breaks.py
file (28 tests) now passes with exact eager/compiled parity across
eager, aot_eager, and inductor backends, and the full CI unit-test
command (pytest --maxfail=1 ./tests/unit_tests/ ... with the existing
deselects) passes 1379 passed / 0 failed.
@WangLingxun
WangLingxun force-pushed the chore/ci-bump-primus-v26.4 branch from de44526 to f9fc67c Compare July 25, 2026 15:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant